BMC Genomics
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match BMC Genomics's content profile, based on 406 papers previously published here. The average preprint has a 0.29% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Guerra-Almeida, D.; Nery, M. F.; Mucherino-Munoz, J. J.; Pessoa-Costa, E.; Tschoeke, D. A.; Nunes-da-Fonseca, R.
Show abstract
Small open reading frames (small ORFs/smORFs/sORFs) were historically overlooked due to their short length (<100 codons) and detection challenges. It is widely accepted that small ORFs are rarely conserved across distant taxa, particularly in deeply rooted ancestors. Growing evidence supports their biological relevance and potential roles in gene innovation. Here, we integrated genomic, transcriptomic, and proteomic datasets to identify and characterize 454 putative small ORFs in Tribolium castaneum. Conservation analyses revealed 230 small ORFs conserved within Coleoptera, 167 in insects, 85 in arthropods, 39 in eukaryotes, and six in bacteria. Conserved small ORFs displayed higher GC content ([~]44.87%), with bacterial-conserved sequences exceeding 50%. Functional enrichment revealed dozens of small ORFs associated with catalytic, transporter, structural, and regulatory roles. Binding site predictions indicated 254 small ORFs possess protein-protein interaction potential. Expression profiles across 14 RNA-seq libraries spanning multiple developmental stages showed 40.1% of small ORFs expressed (TPM >1) in at least one sample, and 16.5% expressed in five or more libraries. Conserved small ORFs showed broader and more stable expression, while lineage-specific ones displayed restricted expression, mostly in gonads or embryos. Isoformic and small CDSs were the most widely expressed classes. Notably, 29 small ORFs exhibited transcriptional and translational support (Swiss-Prot and/or Ribo-seq), 28 being small CDSs. A total of 390 small ORFs had potential paralogs, supporting gene duplication as a mechanism for small ORF emergence. Altogether, our findings reinforce that small ORFs, contrary to initial assumptions, can be deeply conserved and potentially perform key roles in evolutionary processes.
Balasubramanian, R. N.; Umen, J. G.
Show abstract
Cell type specialization is a hallmark of complex multicellular organisms and is usually established through implementation of cell-type-specific gene expression programs. The multicellular green alga Volvox carteri has just two cell types, germ and soma, that have previously been shown to have very different transcriptome compositions which reflect differences in their respective forms and functions. Here we interrogated another potential mechanism for differentiation in V. carteri, cell type specific alternative transcript isoforms (CTSAI). We used pre-existing predictions of alternative transcripts and de novo transcript assembly to compile a list of 1978 loci with two or more transcript isoforms, 67 of which also showed cell type isoform expression biases. Manual curation identified 15 strong candidates for CTSAI, three of which were experimentally verified and provide insight into potential functional differentiation of encoded protein isoforms. Alternative transcript isoforms are also found in a unicellular relative of V. carteri, Chlamydomonas reinhardtii, but there was little overlap in orthologous gene pairs in the two species which both exhibited CTSAI, suggesting that CTSAI observed in V. carteri arose after the two lineages diverged. CTSAIs in metazoans are often generated through alternative pre-mRNA processing mediated by RNA binding proteins (RBPs). We interrogated cell type expression patterns of 126 V. carteri predicted RBP encoding genes and found 40 that showed either somatic or germ cell expression bias. These RBPs are potential mediators of CTSAI in V. carteri and suggest possible pre-adaptation for cell type specific RNA processing and a potential path for generating CTSAI in the early ancestors of metazoans and plants.
Beiki, H.; Murdoch, B. M.; Park, C. A.; Kern, C.; Kontechy, D.; Becker, G.; Rincon, G.; Jiang, H.; Zhou, H.; Thorne, J.; Koltes, J. E.; Michal, J. J.; Davenport, K.; Rijnkels, M.; Ross, P. J.; Hu, R.; Corum, S.; McKay, S.; Smith, T. P. L.; Liu, W.; Ma, W.; Zhang, X.; Xu, X.; Han, X.; Jiang, Z.; Hu, Z.-L.; Reecy, J. M.
Show abstract
Functional annotation of the bovine genome was performed by characterizing the spectrum of RNA transcription using a multi-omics approach, combining long- and short-read transcript sequencing and orthogonal data to identify promoters and enhancers and to determine boundaries of open chromatin. A total number of 171,985 unique transcripts (50% protein-coding) representing 35,150 unique genes (64% protein-coding) were identified across tissues. Among them, 159,033 transcripts (92% of the total) were structurally validated by independent datasets such as PacBio Iso-seq, ONT-seq, de novo assembled transcripts from RNA-seq, or Ensembl and NCBI gene sets. In addition, all transcripts were supported by extensive independent data from different technologies such as WTTS-seq, RAMPAGE, ChIP-seq, and ATAC-seq. A large proportion of identified transcripts (69%) were novel, of which 87% were produced by known genes and 13% by novel genes. A median of two 5 untranslated regions was detected per gene, an increase from Ensembl and NCBI annotations (single). Around 50% of protein-coding genes in each tissue were bifunctional and transcribed both coding and noncoding isoforms. Furthermore, we identified 3,744 genes that functioned as non-coding genes in fetal tissues, but as protein coding genes in adult tissues. Our new bovine genome annotation extended more than 11,000 known gene borders compared to Ensembl or NCBI annotations. The resulting bovine transcriptome was integrated with publicly available QTL data to study tissue-tissue interconnection involved in different traits and construct the first bovine trait similarity network. These validated results show significant improvement over current bovine genome annotations.
Amineni, V. P. S.; Ramapuram, S.; Panfilio, K. A.
Show abstract
BackgroundHalyomorpha halys (brown marmorated stink bug) is an invasive polyphagous pest causing significant agricultural damage worldwide and is an emerging target for RNAi-based pest management. Despite growing interest in dsRNA-based biocontrol, progress is constrained by the lack of tissue-resolved transcriptomic resources covering key biological processes such as feeding, detoxification, and reproduction. Furthermore, our understanding of how RNAi machinery expression varies across tissues remains limited, which impairs both target gene selection and predictions of RNAi efficacy. Critically, the transcriptional response of H. halys to haemolymph-delivered non-specific dsRNA represents a key knowledge gap for evaluating potential non-target immune reactions of dsRNA-based approaches. ResultsField-collected adult males were injected with either nuclease-free water or dsRNA targeting GFP (dsGFP), and transcriptomes were generated from the brain, midgut, salivary glands, and testes. Sequencing produced high-quality datasets with clear tissue-level separation and tight clustering of biological replicates. As expected in targeting a non-endogenous gene, differential expression analysis revealed a limited transcriptional response to dsGFP. Baseline profiling of RNAi pathway genes in controls showed broad expression of core siRNA and miRNA components across all tissues, yet with marked specialisation: two additional Argonaute-2 isoforms and multiple piRNA factors were testes-specific, whereas salivary glands showed strong, restricted expression of nuclease-encoding genes, including a T2 ribonuclease and a non-specific endonuclease. Expression atlases also revealed pronounced tissue partitioning for other protein families. Consistent with their respective functions, secreted trypsins and chymotrypsins are salivary-enriched while the cathepsins for intracellular protein catabolism are midgut-enriched, with brain-centred neuropeptide expression. However, we also uncovered unexpected nuance, such as closely related subfamilies of Cytochrome P450s, which generally function as detoxification enzymes, being partitioned between the midgut, brain or testes. ConclusionsThis work delivers the first tissue-resolved transcriptomic atlas of adult male H. halys, providing a high-resolution resource on compartmentalization of proteolysis, detoxification, and neuroendocrine signalling, as well as for candidate gene discovery in RNAi-based pest control. The modest, tissue-restricted transcriptional response to non-specific dsRNA, together with strong tissue-specific enrichment of some components, offers mechanistic insight into tissue-dependent RNAi efficiency and supports rational dsRNA target selection in H. halys.
Okafor, A.; Adam, Y.; Brors, B.; Adebiyi, E.
Show abstract
BackgroundThe life cycle of Plasmodium parasites is intricate and multistage, alternating between dynamic environments. Temporal regulation of transcription by stage-specific transcription factor binding at particular regulatory regions within gene promoters facilitates its progression. As a result, each new developmental stage is endowed with its unique gene sets, whose just-in-time expression enables the parasite to completely adapt to the necessary circumstances. Our understanding of these transcriptome-level regulatory processes is limited, and more so, a thorough examination of the entire life cycle in the experimentally tractable rodent model organism P. berghei is lacking. ResultsWe performed a genome-wide analysis of RNA-Seq data from different developmental stages of P. berghei. Integrated data from the human malaria parasites P. falciparum and P. vivax demonstrated that Plasmodium parasites have a unique transcriptional signature. We identified the sets of genes differentially expressed at each stage, clustered them based on similarities of their expression profiles, and predicted the regulatory motifs governing their expression. We interpreted the motifs using known binding sites for established eukaryotic transcription factors, including those of the ApiAP2s, and identified eight potentially novel motifs. Additionally, we expanded the annotation of another motif--AGGTAA--found in genes exclusive to erythrocytic development and identified members of the PfMORC and GCN5 complexes among its possible interacting proteins. ConclusionThis study provides new insights into gene usage and its regulation during P. berghei development.
Ahmad, A.; mustafa, h.; Khan, W. A.; Manan, A.; Anwer, I.; Akram, W.
Show abstract
Linkage disequilibrium (LD) and haplotype block structure govern the resolution and utility of genomic selection, marker-assisted selection, and genome-wide association studies (GWAS) in livestock. We performed a comprehensive genome-wide characterization of LD decay, haplotype block architecture, and population diversity across all 24 autosomes in Nili-Ravi buffalo (Bubalus bubalis; n = 85), using 43,543 post-quality-control SNPs. Mean genome-wide r2 was 0.124 (median 0.074) and mean D was 0.540 (median 0.481), with LD half-decay at {approx}70 kb. A total of 133 haplotype blocks encompassing 721 SNPs were identified (Gabriel et al., 2002). Haploview analysis of nine chromosomes harbouring bTB resistance candidate genes revealed contrasting selection signatures: directional selection at innate immune loci (IFNG, TLR1; H < 0.55) versus balancing selection at adaptive immune loci (BoLA-DRB3, SP110; H > 1.0). Critically, BBU15 Block 3 (28.6 kb; OR52E5/NCR1 locus, 47.16 Mb) showed a genome-wide significant integrated haplotype score (iHS; -log1 0 p = 5.408), directly co-localising with the published bTB susceptibility QTL (Bermingham et al., 2014). The TAA haplotype (frequency 53.3%) at this block represents a candidate resistance-associated haplotype for marker-assisted selection. These findings provide essential parameters for SNP panel design and bTB resistance breeding in South Asian buffalo.
Sabino, D. T.; Hortas, A. B.; Villamayor, P. R.; Rasines, I.; Martin, I.; Bouza, C.; Robledo, D.; Martinez, P.
Show abstract
Senegalese sole is a promising European aquaculture species whose main challenge is that captive-born males (F1) are unable to reproduce in farms, hindering breeding programs. Chemical communication through the olfactory system is hypothesized to stem this issue. Although significant advancement in genomic resources has been made over the past decade, scarce information exists on the genomic basis of olfaction, a special sensory system for demersal species like flatfish, which could play a prominent role in reproduction, social and environmental interactions. A full-length transcriptome of the olfactory rosettes including females, males, juveniles and adults, of both F1 and wild origins, was generated at the isoform-level by combining Oxford Nanopore long-read and Illumina short-read sequencing technologies. A total of 20,670 transcripts actively expressed were identified: 13,941 were known transcripts, 5,758 were novel transcripts from known genes, and 971 were novel genes encoding novel transcripts. Special attention was paid to the olfactory receptor gene families (OlfC, OR, ORA and TAAR) expression. Our comprehensive olfactory transcriptome of Senegalese sole provides a foundation for delving into the functional basis of this complex organ in teleost and flatfish. Furthermore, it provides a valuable resource for addressing reproductive management challenges in Senegalese sole aquaculture.
Terefe, E.; Belay, G.; Tijjani, A.; Han, J.; Salim, B.; Hanotte, O.
Show abstract
African zebu cattle (Bos indicus) exhibit remarkable adaptation to extreme thermal conditions, yet the genomic bases of this resilience are not fully elucidated. Ethiopia provides a unique natural setting where closely related zebu populations have divergently adapted to hot-arid (DHETZ) and hot-humid (HHETZ) climates. In this study, we performed whole-genome sequencing of 46 Ethiopian zebu cattle from five populations and compared them with Asian zebu, Sudanese zebu, African taurine, and European taurine breeds. By integrating genome-wide SNP analysis, population genetic structure assessment, and multiple selection scans (iHS, Hp, XP-EHH, and XP-CLR), we identified distinct and shared selection signatures between DHETZ and HHETZ cattle. Ethiopian zebu closely clustered with Sudanese zebu but showed clear divergence from Asian zebu and taurine breeds. Although DHETZ and HHETZ cattle exhibited minimal genetic differentiation, reflecting their shared ancestry, each group displayed unique selection signals. DHETZ cattle showed strong selection in genes involved in oxidative stress regulation, protein folding, mitochondrial function, and vascular remodeling (e.g., SESN2, DNAJC8, GRPEL2, ABLIM3, and AFAP1L1). In contrast, HHETZ cattle displayed signatures in genes associated with immune responses, energy metabolism, and angiogenesis inhibition (e.g., MYD88, PRKACA, PRKACB, and WIF1). Several genes, including VEGFC, TNIP3, and DMXL2, were under selection in both groups, suggesting conserved mechanisms of thermotolerance and reproductive adaptation. These findings reveal a dual pattern of genomic adaptation: while core heat-response pathways are shared, population-specific signatures reflect distinct metabolic and vascular strategies for coping with arid versus humid heat stress. This study provides novel insights into the genomic architecture of environmental adaptation in tropical cattle and offers valuable markers for breeding climate-resilient livestock.
Jayaraman, S.; Chitneedi, P. K.; Kadri, N. K.; Costa-Monteiro-Moreira, G.; Salavati, M.; Charlier, C.; Boichard, D.; Sanchez, M.-P.; Pausch, H.; Kuehn, C.; Prendergast, J. G.; Clark, E. L.
Show abstract
Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.
Azam, S.; Sahu, A.; Pandey, N. K.; Neupane, M.; Tassel, C. P. V.; Rosen, B. D.; Gandham, R. K.; Rath, S. N.; Majumdar, S. S.
Show abstract
BackgroundIndia, with the worlds largest cattle population and more than 50 registered breeds of Bos indicus, stands as a vital reservoir of genetic diversity. However, the abundant diversity among Indian cattle breeds highlights the inadequacy of a single reference sequence to represent the entire genomic content of desi cattle. We recognize the need to capture the genomic differences within the Bos indicus population as a whole, and specifically within the dairy cattle subset by identifying non-reference sequences and constructing a pangenome. FindingFive representative genomes of prominent dairy breeds, including Gir, Kankrej, Tharparkar, Sahiwal, and Red Sindhi, were sequenced using 10X Genomics Linked-Read technology. Assemblies generated from these linked-reads ranged from 2.70 Gb to 2.77 Gb,comparable to the Bos indicus Brahman reference genome. A pangenome of Bos indicus cattle was constructed by comparing the newly assembled genomes with the reference using alignment and graph-based methods, revealing 8 Mb and 17.7 Mb of novel sequence respectively. A confident set of 6,844 Non-Reference Unique Insertions (NUIs) spanning 7.57 Mbs was identified through both methods, representing the pangenome of Indian Bos indicus breeds. Comparative analysis with previously published pangenomes unveiled 2.8 Mb (37%) commonality with the Chinese indicine pangenome and only 1% commonality with the Bos taurus pangenome. Among these, 2,312 NUIs - encompassing [~]2 Mb, were commonly found in 98 samples of the 5 breeds and designated as Bos indicus Common Insertions (BICIs) in the population. Furthermore, 926 BICIs were identified within 682 protein-coding genes, 54 long non-coding RNAs (LncRNA), and 18 pseudogenes. These protein-coding genes were enriched for functions such as chemical synaptic transmission, cell junction organization, cell-cell adhesion, and cell morphogenesis. The protein-coding genes were found in various prominent Quantitative Trait Loci (QTL) regions, suggesting potential roles of BICIs in traits related to milk production, reproduction, exterior, health, meat, and carcass. Notably, 63.21% of the bases within the BICIs call set contained interspersed repeats, predominantly LINEs. Additionally,70.28% of BICIs are shared with other domesticated and wild species, highlighting their evolutionary significance. ConclusionThis is the first report unveiling a robust set of NUIs defining the pangenome of Bos indicus breeds of India. The analyses contribute valuable insights into the genomic landscape of desi cattle breeds.
Longo, A.; Babbucci, M.; Jiao, Z.; Ferraresso, S.; Franch, R.; Bortoletti, M.; Bertotto, D.; Faggion, S.; Ilsley, G. R.; Papadogiannis, V.; Manousaki, T.; Kristoffersen, J.; Tsigenopoulos, C. S.; Macqueen, D. J.; Bargelloni, L.
Show abstract
BackgroundUnderstanding the role of non-coding genomic variation in speciation remains a major challenge in evolutionary biology. Here, we investigated whether regulatory elements contribute to this process between Atlantic and Mediterranean lineages of European sea bass (Dicentrarchus labrax), a well-characterized case-study near speciation where barriers to introgression exist in the presence of connectivity between diverging populations. ResultsWe generated a novel, highly contiguous genome assembly, which was annotated at the epigenomic level using ATAC-seq and ChIP-seq with six embryonic developmental stages and five tissue types in adult fish, identifying thousands of promoters, enhancers, and open chromatin regions. Integrating this annotation with whole-genome sequence data from 65 individuals across three geographically distinct populations, we identified 57,505 outlier SNPs and 332 structural variants (SVs) showing elevated differentiation between Atlantic and East Mediterranean lineages. Outlier SVs affected key regulatory elements and coding genes, while outlier SNPs were enriched in regulatory elements, particularly enhancers active in adult tissues. Local genomic divergence correlated positively with regulatory element density, especially on chromosomes 1, 9, and 18, which are enriched in genes related to osmoregulation, immune response, and oxidative stress -- processes relevant to adaptation across contrasting marine environments. ConclusionsThese findings support a major role for regulatory variation in driving deep lineage divergence through local adaptation.
Ripley, N. E.; Ripley, B.; Li, K.; Hudson, E. E.; Howe, D. K.; Kalbfleisch, T.; Smith, M.; Nielsen, M. K.
Show abstract
BackgroundStrongylus vulgaris is a highly pathogenic equine strongyle whose larval stages migrate through the mesenteric arterial system, yet the molecular basis of its development, host association, and evolutionary biology remains poorly resolved. We generated an integrated genomic and transcriptomic resource to characterize genome structure, gene annotation, stage-associated expression, candidate secreted proteins, isoform diversity, gene-family evolution, putative horizontal gene transfer, and drug resistance-associated homologs in S. vulgaris. ResultsPacBio HiFi sequencing produced a 329.0 Mb genome assembly comprising 3,913 contigs, with a contig N50 of 147 kb and 94.8% BUSCO completeness, substantially improving the prior fragmented draft. Repeat annotation identified 127.5 Mb of repetitive sequence, representing 38.72% of the genome and dominated by unclassified repeats. Integration of RNA-seq-guided annotation with PacBio Iso-Seq evidence refined 16,387 loci and 24,546 transcripts, generating an isoform-retaining discovery proteome of 23,363 predicted proteins. Functional annotation supported 93.5% of predicted proteins and identified 1,530 unknown or weakly annotated candidates. Consensus secretome prediction identified 2,210 high-confidence putative secreted proteins. Stage-associated transcriptomics showed that development was the dominant axis of expression variation, with 5,299 genes differentially expressed between larvae and adults and strong adult sex-associated divergence. Isoform analysis identified 490 high-confidence isoform switches, concentrated primarily in the ML5 female-to-adult female transition. Comparative genomics identified 27,641 orthogroups, 201 S. vulgaris-specific orthogroups, and, after repeat-aware filtering, 48 expanded and 250 contracted gene families. Structure-guided annotation prioritized migratory-stage-enriched secreted candidates, including Cysteine-rich secretory proteins, Antigen 5, and Pathogenesis-related 1 (CAP), Sperm-coating protein (SCP) Tpx-1/Ag5/PR-1/Sc7 protein superfamily (TAPS) -like, von Willebrand factor type A (VWA) - domain, lipid-binding-like, DNase II-like, and peptidase-like proteins. Conservative screening retained 17 putative horizontal gene transfer (HGT) candidates, and drug resistance-homolog analysis recovered 18 nonredundant S. vulgaris homologs without canonical {beta}-tubulin benzimidazole-resistance substitutions. ConclusionsThese results establish the first integrated, high-quality molecular framework for S. vulgaris and show that its parasitic biology is developmentally structured, isoform-rich, and shaped by both conserved strongylid features and lineage-specific gene-family change. This resource provides a foundation for future studies of interhost migration, host interaction, parasite evolution, and genomic surveillance.
Wang, Y.; Nugroho, T.; Johnson, T. J. J.; Couldrey, C.; Harris, B. L.
Show abstract
In recent years, genetic studies have made significant progress in identifying single-nucleotide polymorphisms (SNPs) associated with cattle health and production traits. However, it is still challenging to identify and validate more complicated forms of variation, such as copy number variation (CNV) and other types of structural variation (SV). In this study, SV regions were identified using 37 New Zealand dairy cattle with linked-read sequence data. A transmission-based framework was used to validate these variants at the population scale. 62,438 putative autosomal SV regions were identified with the LongRanger pipeline following the 10x Genomics recommendations. Copy number states for these regions were subsequently estimated via a read-depth based genotyping method using CNVpytor in a population-representative cohort of 2306 animals using Illumina short-read sequencing technology. Mendelian inheritance of copy number states was assessed using linear mixed models incorporating pedigree information, and transmission levels were used to quantify the biological validity of each CNV region. Transmission levels ranged widely, with a mean of 0.5162 across all regions, where higher transmission levels were proportionally enriched for larger SVs. A total of 7218 CNV regions exhibited high transmission levels (>0.9), indicating strong evidence of inheritance. Among these, 7136 overlapped CNV regions reported in one or more public datasets, while 82 high-confidence regions represent previously unreported variants. High-transmission CNV regions tended to show clear, discrete inheritance patterns in trio families, providing the biological evidence that these CNVs are inherited within the population. Together, these results demonstrate that integrating linked-read sequencing with population-scale transmission-based validation provides a robust framework for identifying high-confidence CNV regions. This catalogue of validated CNV regions represents an important resource for downstream functional analyses and the incorporation of structural variation into genomic selection and breeding programs.
Peng, S.; Dahlgren, A. R.; Hales, E. N.; Barber, A. M.; Kalbefleich, T.; Petersen, J. L.; Bellone, R. R.; Mackowski, M.; Cappelli, K.; Capomaccio, S.; Coleman, S. J.; Distl, O.; Giulotto, E.; Waud, B.; Hamilton, N. A.; Leeb, T.; Lindgren, G.; Lyons, L. A.; McCue, M.; MacLeod, J. N.; Metzger, J.; Mickelson, J. R.; Murphy, B. A.; Orlando, L.; Penedo, C.; Raudsepp, T.; Strand, E.; Tozaki, T.; Trachsel, D. S.; Velie, B. D.; Wade, C. M.; Cieslak, J.; Finno, C. J.
Show abstract
1A high-quality reference genome assembly, a biobank of diverse equine tissues from the Functional Annotation of the Animal Genome (FAANG) initiative, and incorporation of long-read sequencing technologies, have enabled efforts to build a comprehensive and tissue-specific equine transcriptome. The equine FAANG transcriptome reported here provides up to 45% improvement in transcriptome completeness across tissue types when compared to either RefSeq or Ensembl transcriptomes. This transcriptome also provides major improvements in the identification of alternatively spliced isoforms, novel noncoding genes, and 3 transcription termination site (TTS) annotations. The equine FAANG transcriptome will empower future functional studies of important equine traits while providing future opportunities to identify allele-specific expression and differentially expressed genes across tissues.
Robic, A.; Hadlich, H.; Costa Monteiro Moreira, G.; Clark, E. L.; Plastow, G.; Charlier, C.; Kuehn, C.
Show abstract
The aim of this study was to compare the circular transcriptome of divergent tissues in order to understand: i) the presence of circular RNAs (circRNAs) that are not exonic circRNAs, i.e. originated from backsplicing involving known exons and, ii) the origin of artificial circRNA (artif_circRNA), i.e. circRNA not generated in-vivo. CircRNA identification is mostly an in-silico process, and the analysis of data from the BovReg project (https://www.bovreg.eu/) provided an opportunity to explore new ways to identify reliable circRNAs. By considering 117 tissue samples, we characterized 23,926 exonic circRNAs, 337 circRNAs from 273 introns (191 ciRNAs, 146 intron circles), 108 circRNAs from small non-coding genes and nearly 36.6K circRNAs classified as other_circRNAs. We suggested in-vivo copying of specific exonic circRNAs by an RNA-dependent RNA polymerase (RdRP) to explain the 20 identified circRNAs with reverse-complement exons. Furthermore, for 63 of those samples we analyzed in parallel data from total-RNAseq (ribosomal RNAs depleted prior to library preparation) with paired mRNAseq (library prepared with poly(A)-selected RNAs). The high number of circRNAs detected in mRNAseq, and the significant number of novel circRNAs, mainly other_circRNAs, led us to consider all circRNAs detected in mRNAseq as artificial. This study provided evidence that there were 189 false entries in the list of exonic circRNAs: 103 artif_circRNAs identified through comparison of total-RNAseq/mRNAseq using two circRNA tools, 26 probable artif_circRNAs, and 65 identified through deep annotation analysis. This study demonstrates the effectiveness of a panel of highly expressed exonic circRNAs (5-8%) in analyzing the diversity of the bovine circular transcriptome.
KRICK, M. V.; DESMARAIS, E.; SAMARAS, A.; GUERET, E.; DIMITROGLOU, A.; PAVLIDIS, M.; TSIGENOPOULOS, C. S.; GUINAND, B.
Show abstract
BackgroundWhile the stress response inspired genome-wide epigenetic studies in vertebrate models, it remains mostly ignored in fish. We modified the epiGBS (epiGenotyping By sequencing) technique to explore changes in genome-wide cytosine methylation to a repeated acute stress challenge in the nucleated red blood cells (RBCs) of the European sea bass (Dicentrarchus labrax). This species is widely studied in both the natural and farmed environments, including issues regarding health and welfare. ResultsWe retrieved 501,108,033 sequencing reads after trimming, with a mean mapping efficiency of 73.0% (unique best hits). Fifty-seven differentially methylated cytosines (DMCs) close to 51 distinct stress-related genes distributed on 17 of 24 linkage groups (LGs) were detected between RBCs of pre- and post-stress individuals. Literature surveys indicated that thirty-eight of these genes were previously reported as differentially expressed in the brain of zebrafish, most of them involved in stress coping differences. DMC-related genes associated to the Brain Derived Neurotrophic Factor, a protein that favors stress adaptation and fear memory, are especially relevant. ConclusionWe provide an improved epiGBS protocol with increased multiplexing and sequencing capacities that offer new opportunities to improve data acquisition and to investigate important biological processes at a genome-wide level, such as the stress response. Minimally invasive RBCs deserve more attention to investigate the epigenetic response to stress without sacrificing fish.
Poutougnigni Matenchi, Y.; Matthew, H.
Show abstract
BackgroundGudali, a West and Central African shorthorn zebu renowned for its dual-purpose potential, is a key genetic resource in regional livestock production. It has recently been used in major crossbreeding programs, notably with Italian Simmental, to produce the Simgud hybrid. These initiatives aim to combine the exceptional adaptive traits of Gudali with the superior productive performance of Simmental. However, the genomic impact of such crossbreeding on both adaptation and performance remains poorly understood. In this study, we investigated candidate signatures of selection and their associations with quantitative trait loci (QTL) and functionally important genes in the genomes of Gudali and Simgud. Our findings provide insights to guide reasoned, targeted breeding strategies that enhance productivity in tropical environments while preserving adaptive potential. ResultsFrom a dataset of 539 Gudali and 139 Simgud genotyped with the GeneSeek GGP+ Bovine 100K array, we performed a two-step imputation to whole genome and used the resulting dataset to detect candidate selection signatures using Tajimas D, the integrated haplotype score (iHS), the fixation index (FST), and the cross-population extended haplotype homozygosity (XP-EHH). Combining the identified regions under selection, together with gene expression and quantitative trait loci (QTL) databases, we further investigated the genomic targets of natural and artificial selection to identify functional candidate genes underlying adaptation mechanisms. In general, the regions under selection were associated mainly with immunity, food scarcity, thermotolerance and various production traits as important selection targets. For instance, signals detected on BTA5 and BTA7 shared between Gudali and Simgud harbored many olfactory genes (OR2O2, OR7A94, OR7H5P) and taste receptors in Simgud (TAS2R42 and TAS2R46) essential to detect forages and predators in grazing lands. Analyzing 27 tissues, we found that the genes within the regions under selection were mostly enriched for those overexpressed in testis, lung, kidney and hypothalamus. ConclusionBy integrating signatures of selection with information from QTL and gene expression, we identified four genes whose relevance was supported not only by selection signals but by additional functional evidence. For instance, GAB2 for response to trypanosome infection and EYA1 associated with heat/drought adaptation, both needed to thrive in challenging tropical environments.
Dewari, P. S.; Regan, T.; Chapuis, A. F.; Florea, A.; Furniss, J. J.; Clark, T. C.; Taylor, R. S.; Bean, T. P.
Show abstract
BackgroundThe Pacific oyster (Crassostrea/Magallana gigas) is increasingly recognised as a model marine invertebrate. Valued for both ecological and commercial importance, Pacific oysters are farmed widely, supporting global food security by providing a sustainable nutrient-rich source of protein. Despite the significant and recurring economic losses caused by Ostreid herpesvirus (OsHV-1) outbreaks, only a limited number of studies have examined host-pathogen interplay at single-cell resolution. The few available studies largely focus on circulating immune cells (haemocytes), thereby overlooking the complexity of host responses across different tissues and organs. ResultsWe present a detailed single-nucleus transcriptomic atlas of the whole Pacific oysters, including during OsHV-1 infection. A total of 18 distinct transcriptomic clusters were resolved, capturing major cell populations from the gill, mantle, hepatopancreas, adductor muscle, and haemocytes. Notably, three populations- gill ciliary cells, hepatopancreas cells, and an immune-enriched cluster 1- exhibited pronounced transcriptomic responses to OsHV-1 infection. Across the 6, 24, 72, and 96 hours post-infection (hpi) time course, viral transcripts were detected almost exclusively at 72 hpi, with enrichment primarily in adductor muscle cells and two immune cell populations- immature haemocytes, and hyalinocytes. ConclusionsOur findings suggest potential entry portals and tissue-specific replication sites for the OsHV-1 virus in Pacific oysters. This atlas resource provides a high-resolution cellular framework for understanding host-virus interactions and establishes a foundation for future investigations into herpesvirus pathogenesis in marine invertebrates.
Godia, M.; Hammoud, S. S.; Naval-Sanchez, M.; Ponte, I.; Rodriguez-Gil, J. E.; Sanchez, A.; Clop, A.
Show abstract
BackgroundThe mammalian mature spermatozoon has a unique chromatin structure in which the vast majority of histones are replaced by protamines during spermatogenesis and a small fraction of nucleosomes are retained at specific locations of the genome. The chromatin structure of sperm remains unresolved in most livestock species, including the pig. However, its resolution could provide further light into the identification of the genomic regions related to sperm biology and embryo development and it could also help identifying molecular markers for sperm quality and fertility traits. Here, for the first time in swine, we performed Micrococcal Nuclease coupled with high throughput sequencing on pig sperm and characterized the mono-nucleosomal (MN) and sub-nucleosomal (SN) chromatin fractions. ResultsWe identified 25,293 and 4,239 peaks in the mono-nucleosomal and sub-nucleosomal fractions, covering 0.3% and 0.02% of the porcine genome, respectively. A cross-species comparison of nucleosome-associated DNAs in sperm revealed positional conservation of the nucleosome retention between human and pig. Gene ontology analysis of the genes mapping nearby the mono-nucleosomal peaks and identification of putative transcription factor binding motifs within the mono-nucleosomal peaks showed enrichment for sperm function and embryo development related processes. We found motif enrichment for the transcription factor Znf263, which in humans was suggested to be a key regulator of the genes with paternal preferential expression during early embryo development. Moreover, we found enriched co-occupancy between the RNAs present in pig sperm and the RNA related to sperm quality, and the mono-nucleosomal peaks. We also found preferential co-location between GWAS hits for semen quality in swine and the mono-nucleosomal sites identified in this study. ConclusionsThese results suggest a clear relationship between nucleosome positioning in sperm and sperm and embryo development.
Stiens, J.; Tan, Y. Y.; Joyce, R.; Arnvig, K. B.; Kendall, S. L.; Nobeli, I.
Show abstract
A whole genome co-expression network was created using Mycobacterium tuberculosis transcriptomic data from publicly available RNA-sequencing experiments covering a wide variety of experimental conditions. The network includes expressed regions with no formal annotation, including putative short RNAs and untranslated regions of expressed transcripts, along with the protein-coding genes. These unannotated expressed transcripts were among the best-connected members of the module sub-networks, making up more than half of the hub elements in modules that include protein-coding genes known to be part of regulatory systems involved in stress response and host adaptation. This dataset provides a valuable resource for investigating the role of non-coding RNA, and conserved hypothetical proteins, in transcriptomic remodelling. Based on their connections to genes with known functional groupings and correlations with replicated host conditions, predicted expressed transcripts can be screened as suitable candidates for further experimental validation.